Papers with French dataset

8 papers
A Multimodal French Corpus of Aligned Speech, Text, and Pictogram Sequences for Speech-to-Pictogram Machine Translation (2024.lrec-main)

Copied to clipboard

Challenge: Existing algorithms for the automatic translation of spoken language into pictogram units are lacking for language impairments.
Approach: They propose to use a French dataset that contains 230 hours of speech resources to create a rule-based pictogram grammar with a restricted vocabulary and a discussion of strategic decisions.
Outcome: The proposed model is validated through multiple post-editing phases by expert annotators and is freely available under a non-commercial licence.
BLM-AgrF: A New French Benchmark to Investigate Generalization of Agreement in Neural Networks (2023.eacl-main)

Copied to clipboard

Challenge: Existing benchmarks for deep learning are based on massive amounts of data, which are effective in hiding some of the shallowness of the learned models.
Approach: They propose to use a French dataset to learn the underlying rules of subject-verb agreement in sentences, inspired by visual IQ tests known as Raven’s Progressive Matrices.
Outcome: The proposed method is based on Raven’s Progressive Matrices, a visual IQ test, and a dataset built using the BLM framework.
HISTOIRESMORALES: A French Dataset for Assessing Moral Alignment (2025.naacl-long)

Copied to clipboard

Challenge: HistoiresMorales is a dataset based on moralStories in French . it is based upon annotations of moral values within the dataset .
Approach: They propose a dataset in French that aims to align language models with moral values . they use annotations to ensure their alignment with French norms .
Outcome: The proposed dataset guarantees grammatical accuracy and adaptation to the French cultural context.
Multilingual prediction of Alzheimer’s disease through domain adaptation and concept-based language modelling (N19-1)

Copied to clipboard

Challenge: Existing work on speech and language models has been limited by the size of available datasets.
Approach: They propose to augment a small French dataset with a much larger English dataset to augment the language model to model the order in which information units are produced by dementia patients and controls.
Outcome: The proposed model improves classification performance in English and French separately.
He said “who’s gonna take care of your children when you are at ACL?”: Reported Sexist Acts are Not Sexist (2020.acl-main)

Copied to clipboard

Challenge: Sexism is prejudice or discrimination based on a person's gender.
Approach: They propose to use a French dataset annotated for sexism detection to characterize sexist content and to train deep learning experiments on tweets.
Outcome: The proposed dataset is the first to be used for sexism detection in France and constitutes a first step towards offensive content moderation.
Give me your Intentions, I’ll Predict our Actions: A Two-level Classification of Speech Acts for Crisis Management in Social Media (2022.lrec-1)

Copied to clipboard

Challenge: Using social networks, social media is a vital tool for emergency management and social media has been used to generate valuable information in crisis situations.
Approach: They propose to measure for the first time the role of SA on urgency detection in tweets . they propose to use a two-layer annotation scheme to annotate tweets for both SA and urgency .
Outcome: The proposed scheme combines two-layer annotation scheme and deep learning experiments to detect SA in a crisis corpus.
New Semantic Task for the French Spoken Language Understanding MEDIA Benchmark (2024.lrec-main)

Copied to clipboard

Challenge: Intent classification and slot-filling tasks are essential tasks of Spoken Language Understanding (SLU).
Approach: They propose to use a MEDIA SLU dataset to train a multilingual model to achieve both tasks jointly.
Outcome: The proposed model can be trained on multiple datasets including the MEDIA dataset and extends to more tasks and use cases.
CIS-BWE: Chaos-Informed Speech Bandwidth Extension (2026.acl-long)

Copied to clipboard

Challenge: CIS-BWE introduces two chaos-informed discriminators for capturing the deterministic chaos from speech.
Approach: They propose a novel adversarial Bandwidth Extension framework that introduces two chaos-informed discriminators for capturing the deterministic chaos from speech.
Outcome: The proposed framework achieves better performance across nine subjective and objective evaluation metrics with a 40x reduction in discriminator size and overall 0.5x fewer parameters, establishing a new baseline in the BWE task.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations